Papers by Munmun De Choudhury

9 papers
Responsible Evaluation of AI for Mental Health (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent.
Approach: They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation.
Outcome: The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation.
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform (2025.emnlp-main)

Copied to clipboard

Challenge: 108K drug overdose deaths in 2022, according to NIDA .
Approach: They propose a large-scale study of OUD-related myths on YouTube with clinical experts to validate 8 pervasive myths and release an expert-labeled video dataset.
Outcome: The proposed model reduces annotation time and cost by over 76% compared to experts and full LLM labeling.
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use (2025.naacl-long)

Copied to clipboard

Challenge: Adverse Drug Reactions (ADRs) from psychiatric medications are the leading cause of hospitalizations among mental health patients.
Approach: They propose a benchmark and a framework to evaluate LLMs' ability to detect ADRs . they find that LLM responses are more complex and harder to read than experts .
Outcome: The proposed framework evaluates LLMs' ability to detect and deliver expert-aligned mitigation strategies.
Who’s Asking? Simulating Role-Based Questions for Conversational AI Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Language model users embed personal and social context in their questions.
Approach: They propose a framework for simulating role-based questions using a taxonomy of asker roles for patients, caregivers, practitioners.
Outcome: The proposed framework simulates 15,321 questions that embed each asker role’s goals, behaviors, and experiences.
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored.
Approach: They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies.
Outcome: The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics.
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs’ Self-consistency in Closed Domains Via Adversarial Nudge (2026.acl-long)

Copied to clipboard

Challenge: Claude exhibits strong resilience, while GPT and Grok demonstrate moderate resilience . open models fall short significantly, while proprietary models exhibit weak resilience compared to open models .
Approach: They propose a framework for stress testing factual fidelity in large language models in the presence of adversarial nudges.
Outcome: The proposed model is robust to adversarial nudges in two closed domains.
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations.
Approach: They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings.
Outcome: The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria.
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language.
Approach: They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication.
Outcome: The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue.
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored.
Approach: They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities.
Outcome: The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations